Problem Set 1: Statistics Fundamentals

PLS 206 — Applied Multivariate Modeling

Author

Grey Monroe

Published

January 1, 2026

Note

Due: Monday, October 5, 1:10 PM (start of class) — submitted through Canvas.

Parts 1–8 use the R skills from the Friday crash course. Parts 9 and 10 use the Week 1 lecture. The Week 1 script, Week1_Stats_Fundamentals.R, has the code for every part: Part E of the script covers Parts 1–8 here, sections 10–14 cover Part 9, and sections 21–25 cover Part 10.


Submission Instructions

Upload two files to Canvas:

  1. Written answers — a .pdf document with your written responses, screenshots, and interpretations.
  2. R script — a .R file with all the code you used to produce your answers.

List the first and last names of any collaborators in both files.

Naming convention: HW1emailID.R (code) and HW1emailID.pdf (written answers), where emailID is the part of your UC Davis email before the @.

Example: Grey Monroe would submit HW1gmonroe.R and HW1gmonroe.pdf.


Part 1: Create and Manipulate a Data Frame

Q1. Use ?set.seed or Google to look up the set.seed() function. Explain in your own words what it does and why it matters for reproducible research.

Q2. Create a data frame called df using the code below. Then find the sum of all values in df.

set.seed(1522)   # replace 1522 with your own favorite number
df <- data.frame(
  yield = rnorm(100, mean = 100, sd = 5),
  temp  = rnorm(100, mean = 10,  sd = 1)
)

Q3. Find the mean of yield.

Q4. Find the 8th value of yield.


Part 2: Import and Examine the Wine Dataset

This dataset contains results of a chemical analysis of 3 cultivars of wine from the same region of Italy. Thirteen chemicals were measured for each sample. It was modified from a dataset freely available through the UCI Machine Learning Repository.

Download winedata.csv from the course data folder and read it into R:

wine <- read.csv("winedata.csv", header = TRUE)
str(wine)
Tip

Use str() to inspect your data — it shows variable types and a preview of values, which is more informative than View() for checking your import.

Q5. What are the dimensions of the wine dataset (rows × columns)?

Q6. What type of object is wine?

Q7. Provide the name of one column in wine that contains:

    1. a categorical variable
    1. a quantitative integer variable
    1. a quantitative numeric variable

Part 3: Identify and Fix Common Coding Errors

Each chunk below contains a single bug. Fix the error and provide the corrected code in your R script. Write one sentence describing each bug in your written answer document.

A useful reference: Interpreting Common R Errors

Q8.

DF[1:5, 1:2]   # subset rows 1–5 and columns 1–2 of the data frame from Part 1

Q9.

vector.a  <- c(0.2, 0.5, 0.8, 0.01, 0.03)
vector.b  <- c(12, 7, 5, 4, 14)
dataframe <- rbind(vector.a vector.b)   # combine into a data frame

Q10.

color5 <- wine[wine Color >= 5]   # subset to rows where Color >= 5

Part 4: Calculate Basic Summary Statistics

Q11. Write code to obtain the mean of all quantitative columns in wine. Which column has the highest mean?

Q12. Write code to obtain the standard deviation of all quantitative columns in wine. Which column has the lowest standard deviation?


Part 5: Subset the Dataset

Reduce wine to retain only these columns: Cultivar, Alcohol, Ash, Color, Flav, Mg, and OD. Save this as a new object.

Q13. Use a function to display the top rows of your reduced data frame. Paste a screenshot of the output in your written answer document.


Part 6: Examine Correlations Between Variables

Load ggplot2 and GGally, then use ggpairs() to create a scatter plot matrix for the quantitative variables in your reduced dataset.

library(ggplot2)
library(GGally)

# ggpairs on columns 2 through 7 (the continuous variables)

Q14. Paste a screenshot of your scatter plot matrix in your written answer document.

Q15. Which 3 pairs of variables are most strongly correlated? (Consider the absolute value of correlations — include both positive and negative relationships.)


Part 7: Examine Data by Cultivar

Q16. Write code to obtain the sample size for each of the three cultivars. List the values in your written answer.

Q17. Create a second ggpairs() scatter plot matrix that includes:

  • Data for each cultivar shown in a different color
  • Scatter plots for all pairs of quantitative variables
  • Box plots and density plots
  • Correlations with both the overall value and the per-cultivar value

Paste a screenshot of your plot in your written answer document.

Q18. Write code to find the mean value by cultivar for each of the 6 quantitative chemicals in your reduced dataset. Then examine these means along with the plots to identify:

    1. A variable where the cultivar values appear clearly different
    1. A variable where the cultivar values are mostly overlapping

Part 8: Flav vs. OD

Q19. Using the second scatter plot matrix from Q17, find the correlations between Flav and OD for each cultivar and the overall correlation.

Describe: How do the per-cultivar correlations compare to the overall value? Is the relationship consistent across cultivars, or does one cultivar primarily drive the overall pattern?


Part 9: Uncertainty and Comparing Two Groups (Monday lecture)

Tip

The cultivar names are lowercase in the data (barbera, barolo, gringnolino), and R is case-sensitive: wine$Cultivar == "Barolo" matches nothing. Check with unique(wine$Cultivar).

Q20. For the Barolo wines only, calculate the mean, standard deviation, standard error, and 95% confidence interval of Alcohol by hand (using mean(), sd(), length(), sqrt(), and qt()). Then check your interval against t.test(). Report all five numbers.

Q21. In one or two sentences each:

    1. What is the difference between the SD and the SE you calculated in Q20? Which one describes individual wines, and which one describes your estimate of the mean?
    1. If you had measured four times as many Barolo wines, what would you expect to happen to the SD? To the SE?

Q22. Compare proanthocyanins (Proa) between Barolo and Gringnolino.

    1. Make a plot showing the raw data for the two cultivars (e.g. a boxplot with the points jittered on top). Paste it in your written answer document.
    1. Run a two-sample t-test. Report the difference in means, its 95% confidence interval, the t statistic, and the p-value.
    1. Write one sentence interpreting the result in the units of the data, the way you would in the results section of a paper.

Q23. Calculate a p-value for the same comparison without t.test(), by shuffling the cultivar labels. Use set.seed(), shuffle the labels 10,000 times with sample(), and record the difference in means each time. What fraction of shuffled differences are at least as large (in absolute value) as the observed difference? How does this compare to the p-value from Q22?


Part 10: More Than Two Groups and Effect Sizes (Wednesday lecture)

Q24. Run a one-way ANOVA testing whether Flav differs among all three cultivars. Report the F statistic and p-value. Then use TukeyHSD() to determine which pairs of cultivars differ.

Q25. Check the assumptions of your Q24 model: make a Q–Q plot of the residuals and run a Shapiro–Wilk test on them. Then run the non-parametric alternative, kruskal.test(). Does your conclusion from Q24 change?

Q26. Effect sizes.

    1. Calculate Cohen’s d for your Q22 comparison (Proa, Barolo vs. Gringnolino), and for the same two cultivars on Alcohol.
    1. Calculate eta-squared (SSbetween / SStotal) for the ANOVA in Q24 (Flav), and for a second ANOVA on Mg.
    1. All four of these tests give p < 0.01. Explain in two or three sentences why the p-values alone don’t tell you which differences are biologically large, and what the effect sizes add.